Centralised vs Federated HIE
The choice of where clinical data lives is the decision that shapes every other decision in a health information exchange. It determines the privacy risk profile, the analytics capability, the availability requirements and — usually decisively — whether the institutions holding the data will participate at all.
1. Problem
A clinician sees a patient who has been treated elsewhere. The information needed to treat them safely exists, in another organisation's system. How is it made available, reliably, lawfully, and quickly enough to be useful?
2. Context
Applies when: multiple organisations hold clinical records; patients move between them; and there is a mandate to make records available across organisational boundaries.
The forces in tension:
- Clinicians need information quickly and completely
- Institutions want to retain custody of their data
- Data protection law constrains where identifiable data may be stored
- Public health and planning need population-level analysis
- Connectivity and operational capacity are finite
3. Architecture
Centralised
EMR A ──┐
EMR B ──┤ push on save / scheduled
Lab ──┼────────────▶ ┌──────────────────────┐
CHW ──┘ │ Central shared │◀── query ── Consumers
│ health record │
│ + client registry │──────────▶ Analytics
└──────────────────────┘
Sources send data as it is created. The central store answers all queries.
Federated
┌────────────────────┐
Query ──────────▶│ Record locator │ which sources hold data
│ + client registry │ for this patient?
└─────────┬──────────┘
│ fan-out
┌───────────────┼───────────────┐
▼ ▼ ▼
EMR A EMR B Lab
│ │ │
└───────────────┴───────────────┘
assemble and return
Sources keep their data. The exchange knows only where data exists.
Hybrid
Sources ──push summary──▶ ┌──────────────────────────┐
│ Central: client registry,│◀── fast query
│ record locator, patient │
│ summary (IPS-shaped) │
└────────────┬─────────────┘
│ on-demand detail fetch
┌────────────┴────────────┐
▼ ▼
EMR A Lab
4. Components
| Component | Centralised | Federated | Hybrid |
|---|---|---|---|
| Client registry | Required | Required | Required |
| Facility registry | Required | Required | Required |
| Interoperability layer | Required | Required | Required |
| Central clinical repository | Required | No | Summary only |
| Record locator service | Optional | Required | Required |
| Source-side query API | No | Required, highly available | Required for detail |
| Terminology service | Required | Required | Required |
| Consent service | Required | Required | Required |
The federated model's distinguishing requirement is that every source must operate a highly available query API. That is a commitment each participating institution must make and resource, and it is where federated designs most often fail in practice — not in the central architecture, but at the twentieth hospital that cannot keep its API up.
5. Data flow
Centralised — write path: source saves → interoperability layer resolves identity, translates terminology, validates → central repository stores → audit.
Centralised — read path: consumer queries → authorisation and consent evaluated → central repository responds → audit.
Federated — read path: consumer queries → identity resolved → locator returns sources → parallel queries with per-source authorisation → responses assembled, deduplicated and returned → audit at every hop.
Note the deduplication step. The same laboratory result may be returned by both the laboratory and the hospital that ordered it. Reconciling duplicates across sources at query time, with different identifiers and different codings, is a genuinely hard problem that centralised designs solve once at write time.
6. Advantages
| Centralised | Federated |
|---|---|
| Fast, predictable query latency | Sources retain custody — often the only politically viable option |
| Works when sources are offline | No stale copies; the source is always authoritative |
| Population analytics is direct | Smaller central privacy footprint |
| One place to secure, audit and monitor | Easier to satisfy laws prohibiting central storage |
| Deduplication and normalisation done once | No large-scale data migration to begin |
| Simpler to operate | Institutions can join without surrendering data |
7. Disadvantages
| Centralised | Federated |
|---|---|
| A single very high-value target | Latency bounded by the slowest source |
| Requires legal basis for central storage | Every source must be highly available |
| Data can be stale relative to source | Population analytics is difficult or impossible |
| Institutions may refuse to participate | Distributed debugging is hard |
| Storage and retention costs concentrate | Deduplication at query time is complex |
| Central store becomes politically contested | Consent must be enforced consistently by every participant |
8. When to use
Centralised when: there is a clear central mandate and legal basis; edge connectivity is unreliable; population analytics and surveillance are primary goals; and central operational capacity exists.
Federated when: law or institutional politics prohibit central storage; connectivity between institutions is reliable; sources are capable of operating APIs at the required availability; and the primary use case is point-of-care retrieval rather than analysis.
Hybrid when: neither of the above cleanly applies — which is most of the time. Centralise identity, the index and a summary set; federate the detail.
9. When not to use
Do not choose centralised if the legal basis is unresolved. Building it and seeking authorisation afterwards is how programmes are stopped after delivery.
Do not choose federated if participating institutions cannot commit to availability. A federated query that times out on half its sources returns a partial record with no indication of what is missing — which is clinically worse than no record at all, because it looks complete.
Do not choose either if there is no client registry. Both models depend entirely on identity resolution, and neither degrades gracefully without it.
10. Example technologies
| Role | Options |
|---|---|
| Interoperability layer | OpenHIM, Mirth/NextGen Connect, Apache Camel |
| Central repository | HAPI FHIR, Firely, Aidbox, managed cloud FHIR |
| Client registry | OpenCR, SanteMPI |
| Record locator | IHE XDS registry, or a FHIR-based index |
| Document sharing | IHE XDS/XCA, or FHIR MHD |
| Terminology | Snowstorm, Ontoserver, HAPI terminology |
| Analytics | DHIS2, warehouse or lakehouse |
Making the decision
The questions that actually decide it, in order:
- What does the law permit? This is a legal opinion, obtained in writing, not an architectural preference.
- Will the largest data holders participate? If the three biggest hospitals will not send data centrally, a centralised design has already failed.
- What is the primary use case? Point-of-care retrieval and population analytics pull in different directions.
- What availability can sources actually sustain?
- Who operates the central components, and are they funded beyond the project?
Record the answer and the reasoning as an ADR. This decision will be revisited, and the next team needs to know which constraint drove it — because when that constraint changes, the decision should change too.
References
- OpenHIE architecture — https://ohie.org/
- IHE IT Infrastructure profiles — https://www.ihe.net/
- International Patient Summary IG — https://hl7.org/fhir/uv/ips/
- Health information exchange